Papers by Arlindo Rodrigues Galvão Filho

4 papers
Safety Is Not Universal: The Selective Safety Trap in LLM Alignment (2026.findings-acl)

Copied to clipboard

Challenge: Existing safety evaluations of large language models aggregate harms under generic categories such as "Identity Hate" a bilingual benchmark identifies a selective safety trap, where defense rates vary by up to 42% within the same model solely based on the target group.
Approach: They propose a bilingual adversarial benchmark to audit selective safety in large language models . defense rates vary by up to 42% within the same model solely based on target group .
Outcome: The proposed benchmark identifies a selective safety trap in large language models . defense rates vary by up to 42% within the same model solely based on the target group.
Modeling, Evaluating, and Embodying Personality in LLMs: A Survey (2025.findings-emnlp)

Copied to clipboard

Challenge: This survey provides a comprehensive overview of the LLM-driven personality scenario.
Approach: This survey provides a comprehensive overview of the LLM-driven personality scenario.
Outcome: The proposed taxonomy analyzes the limitations of existing methods and identifies key research gaps.
Proxy Barrier: A Hidden Repeater Layer Defense Against System Prompt Leakage and Jailbreaking (2025.findings-emnlp)

Copied to clipboard

Challenge: Prompt injection and jailbreak attacks remain a critical vulnerability for large language models . a lightweight defense that interposes a proxy LLM between the user and the target model addresses this vulnerability .
Approach: a lightweight proxy LLM is interposed between the user and the target model to prevent prompt injection and jailbreak attacks.
Outcome: ProB outperforms baselines and achieves up to 98.8% defense effectiveness . it is deployable entirely at the API level and requires no access to model weights or prompts .
BRSpeech-DF: A Deep Fake Synthetic Speech Dataset for Portuguese Zero-Shot TTS (2025.emnlp-main)

Copied to clipboard

Challenge: ADD detection is a key area of research for low-resource languages like Portuguese, which lacks high-quality datasets.
Approach: They propose to provide the first publicly available ADD dataset for Portuguese, encompassing both Brazilian and European variants.
Outcome: The proposed dataset contains over 458,000 utterances, including a smaller portion of real speech from 62 speakers and a large collection of synthetic samples generated using multiple zero-shot text-to-speech (TTS) models, each conditioned on the original speaker’s voice.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations